Skip to content

feat: add push-to-talk voice dictation with Parakeet - #240

Merged
Zongwei9888 merged 2 commits into
HKUDS:mainfrom
EduCosta85:feat/composer-voice-dictation-parakeet
Sep 22, 2026
Merged

Zongwei9888 merged 2 commits into
HKUDS:mainfrom
EduCosta85:feat/composer-voice-dictation-parakeet

Conversation

@EduCosta85

Copy link
Copy Markdown
Contributor

Summary

This PR adds push-to-talk voice input directly to the composer. Spoken text is recorded locally, transcribed via an OpenAI-compatible speech-to-text server (such as NVIDIA Parakeet served by mlx-audio), and inserted directly into the prompt draft at the caret position.

What is included

  1. Protocol & App Server:
    • dictation/status method returning endpoint capability (available, model, maxAudioSeconds). Added to READ_METHODS for safe retries.
    • dictation/transcribe method receiving { audio, mimeType, language?, projectId? } and returning { text, model }. Kept out of READ_METHODS.
  2. Configuration (DictationConfig):
    • Declared under dictation in user configuration (~/.deepcode/deepcode_config.json).
    • Validates endpoint (must be absolute HTTP/HTTPS without userinfo, query, or fragment).
    • Project-level configurations cannot configure or hijack the dictation endpoint; repository overrides are dropped during layer sanitization.
  3. Core & Audio Pipeline:
    • Pure container sniffing for audio formats (webm, ogg, mp4, wav, mp3, flac, aac) matching declared MIME types.
    • Caps payloads at 512 KiB decoded / 768 KiB base64 to stay well within the JSON-RPC message envelope.
    • Synchronous SpeechToTextClient targeting POST <endpoint>/audio/transcriptions with multipart form data, no redirect following, and zero response-body leakage in error messages.
  4. Application & Policy:
    • DictationService coordinates egress verification (providers.egress), audio decoding, client dispatch, and error translation.
  5. Desktop & Web UI:
    • useDictation React hook managing recording state, platform-compatible MIME negotiation (MediaRecorder), timers, chunked base64 conversion, and cancellation.
    • Composer microphone button with recording animation, timer hint, discard button (X), and Escape shortcut to cancel.
    • Caret-aware insertion: dictated text is spliced into the existing textarea draft without losing cursor position.
  6. Documentation & Tests:
    • New guide docs/guide/dictation.md covering mlx-audio local Parakeet setup, config reference, and security boundaries.
    • Unit test suites covering audio decoding, client network behaviors, application service, dispatcher RPC, config layering, and frontend hook interactions.

Verification

  • pytest tests/test_dictation.py tests/test_dictation_service.py tests/test_config_layering.py tests/contract/test_protocol_schema.py: 105 passed.
  • ruff check and ruff format --check (v0.15.21): Clean.
  • python -m compileall -q app_server cli core tools workflows: Passed.
  • npm run lint: Clean.
  • npm run typecheck: Passed without errors.
  • npm test -- --run (including useDictation.test.ts): 40 test files passed, 274 tests passed.
  • npm run check:protocol, check:version, check:tauri: All passed.

Support push-to-talk voice dictation directly in the composer.
Speech clips recorded via MediaRecorder are validated, decoded, and
forwarded over JSON-RPC to a local speech-to-text server exposing
an OpenAI-compatible transcription endpoint (such as NVIDIA Parakeet
running on mlx-audio).

- protocol: define dictation/status and dictation/transcribe in JSON-RPC schema
- config: add user-owned DictationConfig with endpoint, model, timeout, and cap settings
- core: implement pure audio container sniffing and OpenAI-compatible client
- application: implement DictationService with model egress policy checks
- desktop: add useDictation hook and composer push-to-talk microphone button
- docs & tests: add complete unit test suites and setup documentation
@Zongwei9888

Copy link
Copy Markdown
Collaborator

Merged into main as 807ff34, with a repair on our side. Thank you @EduCosta85 — this is a substantial, well-structured feature: the config-by-presence design, the "project layer cannot redirect the microphone" rule with its layering tests, the audio sniffing with double payload caps, a client that never echoes response bodies or follows redirects, and the caret-aware insertion in the composer are all as you built them. 105 Python tests plus the 14 hook tests all run offline, which made this reviewable.

What I changed before merging, and why:

  • Scope. core/config.py also introduced AGENT_PHASES, an agents.subagent phase and phase_override() / phase_is_overridden(), with no caller in the PR. Those are a separate feature (a cheaper model tier for sub-agents is on our own backlog), so I removed them here; resolve_phase is as on main. Please do open a follow-up PR for them with a caller and tests — the idea is wanted.
  • Schema vs dispatcher. The dispatcher reads projectId on both methods, but the schema declared dictation/status as EmptyParams and DictationTranscribeParams without projectId. The schema now uses OptionalProjectParams / an optional projectId and the TS was regenerated.
  • Local runner. The docstring described an in-process worker that does not exist, so it now says what the code does (drive the parakeet-mlx CLI). The CLI is located per call rather than at import, the child process gets a minimal environment instead of the harness's provider keys, and the language hint is passed through. Seven tests drive it with a fake CLI. The service also skips the bearer-token lookup for endpoint: "local", which otherwise raised when apiKeyEnv was unset.
  • Client. trust_env=False, so an ambient HTTPS_PROXY cannot re-route audio (and a token) to a host the egress policy never saw.
  • Recorder. audioBitsPerSecond: 32000, so a default 120 s clip stays under the 512 KiB payload cap on browsers whose default bitrate is higher.

One thing neither of us has verified: the exact parakeet-mlx --output-format json --output-dir contract against a real install. If you have it locally and can confirm the output file name, that would close the loop.

Zongwei9888 added a commit that referenced this pull request Sep 22, 2026
Adds an opt-in dictation block (user config only; project layers cannot set it),
dictation/status and dictation/transcribe RPCs, an audio sniffing/decoding
module with payload caps, an OpenAI-compatible speech-to-text client that
never echoes response bodies or follows redirects, a local parakeet-mlx CLI
runner, a Composer microphone with caret-aware insertion, and a guide.
Repair: unrelated phase-routing config changes removed, schema aligned with
the dispatcher, local runner bounded and tested, no proxy trust, recorder
bitrate capped.
Contributed by EduCosta85.
@Zongwei9888
Zongwei9888 merged commit 9b9f07f into HKUDS:main Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants